
1. COMPILING DIFFNORM ON YOUR SYSTEM
====================================
- The source tree uses 'fic' as the main name, as 'fic' encompasses much more than just DiffNorm.
The below mentioned compilation strategies hence resp. deliver 'fic(.exe)' which you yourself 
may decided to rename to 'DiffNorm(.exe)'.

* For GNU/Linux, Mac OS X, ...
	$ ./bootstrap.sh
	$ make -Cbuild install
	- resulting executable can be found in the /bin/ directory

* Compilation on Windows, using Visual C++ 2010 or 2012 or 2013
	- simply open the correct solution (.sln) file, and compile the desired target (Release or Release64)
	- resulting executable can be found in the /src/ directory

2. CONFIGURING DIFFNORM ON YOUR SYSTEM
======================================
- Open `fic.user.conf` and change the following directives
	* basedir : /absolute/path/to/fic/directory
	* datadir : /absolute/path/to/data/directory
	* xpsdir	: /absolute/path/to/experiment/output/directory
	
- On a 32-bits x86 Windows system, try running bin\DiffNorm.exe (or bin\DiffNorm_x86.exe). If 
it raises an error or terminates without any message, you may have to install the Microsoft 
Visual C++ 2013 SP1 Redistributable Package. The x86 version is available for download here:
http://www.microsoft.com/en-us/download/details.aspx?id=40784

- On a 64-bits x64 Windows system, bin\DiffNorm_x64.exe should also work. If it gives an error 
or terminates without any message, you may have to install the Microsoft Visual C++ 2013 
SP1 Redistributable Package. The x64 version is available for download here:
http://www.microsoft.com/en-us/download/details.aspx?id=40784

- On a Linux system, try running bin\DiffNorm. If it does not work, compile your own executable
using the included source (see README).

3. PREPARING YOUR DATABASE FOR DIFFNORM
=======================================
- Databases can be in either [dataDir] or [dataDir]\datasets.

- DiffNorm accepts currently only item data. For now, items should be positive integers.

- DiffNorm has its own database file format. Files in this format generally end with .db. See
the datasets directory for some example databases (mostly from the UCI repository).
- Use `convertdb.conf` to convert your own database to the correct format. The data should be
formatted as the example databases `chess.dat` and `mushroom.dat` (in the datasets directory).
Each row is a transaction, each number is an item that is present in that particular 
transaction. Modify `convertdb.conf` to suit your needs and run it from the command line or
drag-an-drop the config file onto DiffNorm.exe. Commandline:

	C:\DiffNorm\bin> DiffNorm convertdb.conf

- Use analysedb.conf to get some basic statistics about your database, if you like.
The analysis textfile also gives you the translation from the original item numbers to
the numbers as used in the converted database. (Left column: [after]=>[before])

- Check `Data Conversion.txt` for more details.

4. RUNNING DIFFNORM ON YOUR DATA
================================
- Open up `compress.conf`. The most important configuration directives are
	* taskclass = diffnorm
	This specifies which task handler to use. Every project use their own task handler. 
	DiffNorm uses `diffnorm` task handler.

	* command = diffnorm
	This specifies which command to use. 

	* iscName = adult-all-1d
	
	Format: [dbName]-[itemsetType]-[minSup][candidateOrder]	
	dbName			- chess, mushroom, ...
	itemsetType		- all, closed
	minSup			- absolute minimum support level
	candidateOrder		- the standard order as described in Slim paper is 'd'
	
	This directive tells DiffNorm which database and candidates to use. For instance, 
	from the above example, `adult` dataset will be used with minimum support of 1.

	* algo = diffnorm
	To run the diffnorm algorithm. For SLIM/KRIMP, you might have to specify other components too.
	
	* runStrategy = rcpa
	Which strategy to use when running diffnorm. DiffNorm can be run in two different ways: 
		1. Regenerating candidates after testing all 		[rcta]
		2. Regenerating candidates post every acceptance 	[rcpa] <- Recommended

	* genStrategy = gain
	How to sort the candidates? There are two options: gain, usg

	* reEstimateAll = no
	Whether or not to re-estimate of all the candidates in the pool after every candidate acceptance
	
	* pruneStrategy = pop
	Whether or not to prune post acceptance

	* writeReportFile = yes
	Whether or not to write the overall report of the run

	* writeStatsFile = yes
	Whether or not to write the stats file containing all the information about the generated/tested 
	candidates
	
	* reportTranslated = yes
	Whether or not to convert the fic ids back to original fimi ids

- Run diffnorm executable with the `compress.conf`
	* For GNU/Linux, Mac OS X, ...
		$ ./diffnorm compress.conf
	* For Windows, drag-and-drop `compress.conf` into `DiffNorm.exe`

=== 4. Inspect your results ===

- Compression results are stored in the [xpsdir]/diffnorm. Each run gets its own directory based on timestamp.

- In such a directory, you'll find (amongst some unimportant files):
	* diffnorm-[timestamp].conf
	The configuration you used to run this particular experiment.

	* ct-[candidateSet]-[timestamp]-[minsup]-[candNum].ct
	The code table at support level [minsup], which was outputted after processing all the [candNum] 
	candidate item sets. Every row inside `CT_i` section is a code table element for that code table. 
	The pattern is represented by the space-separated items. Between brackets is some meta info: first 
	the usage count of the pattern in that code table current, second the support of the pattern in the 
	database, third the cost, in number of bits, when deleting the itemset from the code table.

	* report.txt
	This file contains all the meta information about the result in general.
	
	* candidate-stats.csv
	This file contains all the candidates generated and tested along side other meta information. Every row
	except the header contains the unique candidate id, estimated gain, actual gain, whether or not the 
	candidate was accepted and estimated usage.
		
- Enjoy! :)
